Frontiers in Bioinformatics
○ Frontiers Media SA
Preprints posted in the last 30 days, ranked by how well they match Frontiers in Bioinformatics's content profile, based on 49 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Li, Q.; Yu, K.
Show abstract
External quality assessment (EQA) of multianalyte assays is commonly interpreted analyte by analyte, although many panels contain known relations among measured features that may reveal joint quality patterns. We propose PathEQA, a feature-graph-guided random forest framework in which a user-supplied graph can represent biochemical pathways, molecular interactions, shared measurement processes, or other domain relations. The same graph is allowed to influence feature representation, node-level candidate generation, and split selection, with an optional local grouped decision. We evaluated the framework in graph-aligned and graph-misspecified simulations and used a six-analyte catecholamine-related liquid chromatography-tandem mass spectrometry EQA data set as an illustrative case study (929 records from 58 laboratories and 117 complete multianalyte panels). In graph-aligned simulations, the grouped variant reduced test root mean squared error by 7.4-9.4% relative to ordinary random forest across training sizes of 60-240, whereas graph misspecification could worsen prediction. In the catecholamine case study, full PathEQA was comparable with ordinary random forest in laboratory-grouped cross-validation (RMSE 0.570 versus 0.569) and modestly better in the final-round temporal holdout (0.307 versus 0.318); a simpler static network-sampling baseline performed best. Dopamine-norepinephrine was the strongest pair, whereas dopamine-norepinephrine-epinephrine best estimated multianalyte failure burden. These results support a general conclusion: feature-graph guidance can improve small-sample multivariate quality assessment when the supplied structure is outcome-relevant, but graph relevance must be tested rather than assumed. Catecholamines serve here as a worked example rather than a restriction of the framework.
Darko, R.; Dwumah, D.; Agyapong, K. S.; Agyenim-Boateng, Y.; Darko Anim, R.; Wisdom Jakper, J.; Owusu-Ansah, N. K.; Owusu-Ansah, R.
Show abstract
Machine learning workflows frequently incorporate data preprocessing to enhance predictive performance. However, the need for Super Learner ensembles made up only of preprocessing-invariant tree-based algorithms remains unexplored. Using three benchmark clinical classification datasets, this study examined how preprocessing affected the Super Learner's prediction performance, learner weight distribution, and oracle behavior. The Heart Disease (207 observations), Indian Liver Patient Dataset (583 observations), and Pima Indians Diabetes (768 observations) datasets were used to create a Super Learner ensemble model that included Classification and Regression Trees (CART), Random Forest, Ranger, and Extreme Gradient Boosting (XGBoost). Models were evaluated under raw and preprocessed data conditions using repeated cross-validation. Predictive performance was assessed using the area under the receiver operating characteristic curve (AUC), Matthews correlation coefficient (MCC), and Brier score. Learner weight allocation and Oracle Gap were compared using paired Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment. Preprocessing produced negligible changes in predictive performance for the Heart Disease and Pima datasets. For the ILPD dataset, preprocessing significantly improved AUC (0.746 to 0.752; adjusted p = 0.0017) and reduced the Brier score (0.177 to 0.175; adjusted p < 0.001). Learner weights remained largely stable, although Random Forest replaced Ranger as the dominant learner for the Heart Disease dataset. Oracle Gaps remained extremely small (<0.002) across all datasets and did not differ significantly between preprocessing conditions. Preprocessing provides limited benefit for Super Learner ensembles composed of preprocessing-invariant learners and does not materially alter their oracle behavior. Preprocessing decisions should therefore be guided by dataset characteristics rather than adopted as a universal modelling practice.
Soto-Garcia, N.; Murillo-Acevedo, N.; Garcia Vinuesa, J.; Islas-Avila, A. L.; D. Davari, M.; Murgas, L.; Hassanin, A.; Orostica, K.; Gonzalez-Puelma, J.; Navarrete, M.; Rebollar-Martinez, A.; Uribe-Paredes, R.; Cadet, F.; Medina-Ortiz, D.
Show abstract
Performance estimates in protein function prediction depend not only on model choice but also on upstream decisions that define the learning problem. Using antioxidant protein classification as a controlled case study, we evaluated how dataset harmonisation, protein representation, redundancy control, and partitioning strategy affect protein machine learning pipelines. We integrated 18,804 records from 12 publicly available dataset entries into a curated consensus dataset of 4,193 protein sequences. One-hot encoding and six pretrained protein language model representations were evaluated as model inputs and as similarity spaces for redundancy reduction and distance-aware splitting. Representation choice substantially altered dataset geometry, retained dataset size, class balance, and downstream evaluation. At representation-specific p90 thresholds, one-hot encoding retained the complete dataset, whereas pretrained embeddings retained between 5% and 25% of sequences. Distance-aware partitioning reduced apparent performance relative to random splitting by up to 0.15 MCC before redundancy control, while this difference narrowed after similarity filtering. Selected configurations nevertheless maintained high performance under stricter evaluation, reaching an MCC of 0.84. These findings show that performance estimates should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.
Show abstract
Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.
Muniz-Chicharro, A.; Tanriver, G.; Gora, A.
Show abstract
Summary: Prot2Surf is a software tool designed for the characterization and prediction of protein association to surfaces. In this application note, Prot2Surf was tested using catalytic domains of the lytic polysaccharide monooxygenases (LPMOs), interacting with native surfaces. The results show that the software can efficiently analyze key binding features, including protein-surface distances, distances between catalytically reactive atoms, and the orientation angle between surface chains and the protein. These features are essential for distinguishing productive binding poses in these protein-surface systems and for understanding interaction patterns that provide guidance on protein engineering. Prot2Surf performs these analyses within seconds to a few minutes, providing a fast and accessible framework to post-process and characterize protein-surface encounter complexes. Availability and implementation: Prot2Surf, which is written in Fortran90, is documented and freely available as open source on GitHub: https://github.com/TUNNELING-GROUP/Prot2Surf. In order to run Prot2Surf, users should also install the SDA software package which is freely available at https://www.h-its.org/downloads/sda7/.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Ab Ghani, N. S.; Matsushita, T.; Noguchi, T.; Kurumida, Y.; Kawada, S.; Ito, T.; Umetsu, M.; Saito, Y.
Show abstract
Motivation Protein language models (PLMs) have emerged as powerful tools for sequence-based prediction of protein function, yet systematic benchmarks comparing frozen embeddings, fine-tuning strategies like Low-Rank Adaptation (LoRA) and classical machine learning (ML) remain limited. We benchmarked four ML strategies: ML using amino acid descriptors (SL-AAFeat), ML using frozen embeddings from 20 PLMs across various pooling strategies (SL-Embed), full model fine-tuning (FT-Full) and LoRA-based fine-tuning (FT-LoRA). Performance was evaluated on the in-house VHH phage display dataset (VHH) for binding affinity prediction and the TAPE fluorescence dataset (FLS and FLS10) for mutational effect prediction. Results Model performance depended strongly on the dataset and adaptation strategy. Max pooling consistently improved embedding-based models, while amino acid descriptors remained competitive under specific datasets and resource constraints. Fine-tuning generally provided the highest predictive performance, but the advantage is not universal. Hyperparameter optimization significantly enhanced FT-LoRA, enabling it to outperform FT-Full on the VHH dataset with less than 10% model parameter adaptation. In contrast, FT-Full achieved the best performance on FLS and FLS10. Several medium-sized PLMs performed comparably to larger models, highlighting favorable performance-efficiency trade-offs. Overall, this paper presents a thorough review of PLM utilization strategies and practical recommendations for selecting suitable strategies based on dataset characteristics and available computational resources. Availability The source code used in this manuscript is available in a Zenodo repository at https://doi.org/10.5281/zenodo.21466255.
Urokov, R.; Khan, A.; Eshboyev, F.; Asadov, D.; Rahman, S.; Kushokova, D.
Show abstract
Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired inference. Of 6,953 atlas BGCs from 182 deduplicated Streptomyces griseus genome accessions, 5,325 silver-labeled BGCs are split into 98 training, 21 validation, and 21 test reference groups. Of the test reference groups, 16 are eligible for retrieval diagnostics. Pfam Jaccard scored Recall@50 of 0.8788, while Pfam-augmented BGC-SetNet scored 0.8472. The combination of ESM and Pfam-augmented BGC-SetNet scored 0.8769. A weighted Pfam Jaccard obtained a slightly higher score of 0.8789, which has a negligible difference compared to unweighted Pfam accard. Our results do not support the claim that sequence-derived representations can recover alternative biosynthetic pathways on this benchmark. Instead, explicit Pfam remains the major signal for this objective. Our results de[fi]ne the curation and pathway-level validation processes that are necessary for a more robust biological test.
Fateh, K.; Yerukala Sathipati, S.
Show abstract
Ovarian cancer is among the deadliest gynecologic malignancies, and its molecular heterogeneity limits accurate prognostic stratification. Although multi-omics approaches have improved predictive modeling, many prioritize predictive performance over biological interpretability, limiting their clinical translation. We developed an interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas. Hierarchical feature selection was combined with a weighted ensemble of ElasticNet, ridge regression, support vector regression, XGBoost, and random forest models to estimate overall survival time in patients with ovarian cancer. Multi-omics integration outperformed every single-modality model, achieving a Pearson correlation of 0.752, a concordance index of 0.779, and a mean absolute error of 8.57 months between estimated and observed survival time, compared with 0.48 for the best single modality. The framework identified a 20-biomarker signature dominated by tumor-associated macrophage and complement genes. In an independent survival analysis, VSIG4 and CD163 remained significant after false discovery rate correction, and the signature raised the concordance index over clinical covariates alone from 0.615 to 0.686Enrichment analysis implicated PI3K-Akt, MAPK, focal adhesion, hypoxia, apoptosis, and p53 signaling pathways. This framework couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.
Hua, X.; Grimaud, G. M.
Show abstract
Accurate enzyme annotation remains a major bottleneck in translating rapidly growing protein sequence data into biological knowledge. Enzyme Commission (EC) prediction is particularly challenging because enzyme functions are organized hierarchically, annotations are often imbalanced across classes, and sequence similarity alone may be insufficient to resolve functional differences. To address these challenges, we developed ESM-ECForest, a two-stage framework that combines protein embeddings generated by the pretrained language model ESM-2 (Evolutionary Scale Modeling 2) with Random Forest classifiers. The first stage distinguishes enzymes from non-enzymes, whereas the second assigns one or more EC numbers to proteins predicted to be enzymatic. On an external benchmark comprising 25,778 protein sequences, ESM-ECForest achieved the highest weighted F1 score among the evaluated methods at all four EC levels, decreasing from 0.94 at Level 1 to 0.90 at Level 4. The largest relative improvements were observed for lyases (EC 4), ligases (EC 6), and translocases (EC 7), although EC 6 and EC 7 remained the most difficult classes internally. Visualization of the ESM-2 embedding space using Uniform Manifold Approximation and Projection (UMAP) revealed clustering patterns consistent with enzyme functional relationships, indicating that biologically relevant information is retained in the pretrained representations prior to supervised classification. These results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation. By combining large-scale sequence representations with a lightweight supervised classifier, ESM-ECForest provides a scalable approach for EC prediction and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.
Kumak, E.; Darde, T.; Konu, O.
Show abstract
Metabolic dysfunction-associated steatotic liver disease (MASLD), the leading cause of chronic liver pathologies worldwide, represents a growing clinical burden. Its diagnosis remains reliant on liver biopsy that limits early detection and the ability to capture molecular changes across disease progression. A systematic understanding of stage-dependent gene expression changes is essential to identify biomarkers and effectively characterize disease mechanisms. Therefore recent studies provided databases for searching genes as well as prediction of multi-gene signatures for disease progression. However, there is still a need for interactive and comprehensive meta-analysis of datasets of MASLD patients with available histological metadata. Herein, we performed a meta-analysis of RNA-seq datasets using NAFLD Activity Score (NAS; n = 897) and fibrosis stage (n = 856) upon conducting pairwise comparisons across histological stages and identified differentially expressed genes associated with disease progression. Most importantly, we provide our findings via a dedicated web server, the MASLD-META NETWORK (https://masld.scilicium.com), enabling users to interactively explore meta-analysis results across diverse network modalities. In addition, we characterized gene expression dynamics across increasing disease stages to identify consistent progression-associated pathways using Louvain clustering. Network-based parameters such as centrality in combination with meta-analysis scores further highlighted central genes and pathways implicated in disease mechanisms. Accordingly, MASLD-META NETWORK enabled an integrative reassessment of recently published gene signatures, identifying COL1A1, COL3A1, THBS2, FBLN5, and PDGFA as the most central genes, and SULF2, MMP14, IL32, GPNMB, and COL3A1 as candidate markers of earlier transcriptional alterations. Network analysis of MASLD associated biological modules further identified LAMA2 and LAMA3 as previously unrecognized central candidate targets.
Zhang, W.; Ji, S.
Show abstract
Background: Bladder cancer has entered an era in which immune checkpoint blockade (ICB) and antibody-drug conjugate (ADC)-based combinations are reshaping clinical management. However, transcriptomic scores that connect prognosis, tumor microenvironment state, and treatment response are incompletely defined. Methods: Open-access TCGA-BLCA RNA-seq, clinical, mutation, copy-number, and RPPA data were downloaded from the Genomic Data Commons (GDC). Tumor-normal differential expressions, survival screening, LASSO-Cox modeling, train-test validation, GEO validation, pathway enrichment, immune signature scoring, mutation/CNV/RPPA support, drug sensitivity prediction, single-cell/spatial localization, and ICB validation were performed using reproducible Python and R scripts. A reduced model was derived using only genes shared by TCGA, GSE13507, and GSE31684. The fixed formula was then applied without refitting to IMvigor210 and GSE176307. Results: A five-gene model composed of EMP1, AHNAK, TNFRSF14, CLEC2D, and GSDMB retained TCGA internal prognostic value (train C-index 0.693, test C-index 0.605, all-sample C-index 0.667; TCGA test log-rank p = 0.015), although GEO survival validation in GSE13507 and GSE31684 was modest. High-risk tumors were enriched for epithelial-mesenchymal transition (EMT), TNF-alpha/NF-kB signaling, inflammatory response, hypoxia, complement, CAF, macrophage, checkpoint, and cytotoxic programs. Single-cell and spatial analyses localized the score to basal tumor, endothelial, fibroblast, and perivascular compartments. In IMvigor210, risk scores were higher in ICB non-responders than responders (Wilcoxon p = 0.044; AUC for non-response = 0.580), high-risk tumors had a lower responder rate (17.6% vs. 28.0%), and high risk predicted poorer OS (log-rank p = 0.016; multivariate continuous risk HR = 3.15, p = 0.044). GSE176307 showed directionally consistent but non-significant response results (AUC = 0.576). Conclusions: The five-gene score is best interpreted not as a standalone universal prognostic classifier, but as a compact stromal-EMT and immune-suppression phenotype associated with inferior ICB response. These findings support a framework linking prognosis, microenvironment biology, immunotherapy resistance, and therapeutic hypotheses in bladder cancer.
de Almeida, D. d. S.; Albuquerque, A. O.; Peixoto Lima, A. M.; Gaieta, E. M.; Souza, J. S.; dos Santos-Costa, A. H.; de Andrade, L. M.; Sampaio, J. V.; Sartori, G. R.; Silva, e. J. H. M. d.
Show abstract
Antibodies generally exhibit high specificity for their cognate epitopes, but structural and physicochemical similarities between distinct epitopes can enable an antibody to recognize different antigens, resulting in cross-reactivity. This property can be exploited for antibody repurposing. To identify epitopes that share such similarities, both sequence- and structure-based approaches can be employed. In this context, 3D Zernike descriptors provide a compact representation of protein surface geometry as numerical feature vectors, enabling quantitative comparisons independently of structural alignment and orientation. Thus, this study aimed to evaluate the application of 3D Zernike descriptors for the structural clustering of antibodies and epitopes and to explore their use in antibody repurposing for the recognition of new targets. To this end, antibody binding sites previously associated with recognition of similar epitopes were analyzed at different structural levels, considering the CDRs, CDRH3, and complete paratopes. Surface similarity was subsequently quantified by calculating the Euclidean distance between their corresponding 3D Zernike feature vectors. Performance was benchmarked against SPACE2. Additionally, different distance thresholds were evaluated based on their ability to recover antibody pairs recognizing the same epitope. The paratope-based approach provided the best balance between the number of identified pairs and precision at a distance threshold of 2.7, whereas epitope clustering showed robust performance up to a distance of 3.0. At these thresholds, the 3D Zernike descriptors identified a greater number of functional pairs than SPACE2 while maintaining comparable precision and identifying complementary sets of antibody pairs.. BTaken together, these findings support the use of 3D Zernike descriptors for structural clustering of antibodies and epitopes and for guiding antibody repurposing G, a highly lethal zoonotic pathogen. Structural screening identified three antibodies with epitopes similar to the NiV target that also showed a consistent binding preference for the target epitope in molecular docking assays. Notably, one candidate, originally directed against a SARS-CoV-2 epitope, formed a stable complex with the NiV epitope, remaining within the 5 [A] RMSD threshold during heated molecular dynamics simulations and emerging as a potential cross-reactive candidate.These results support the use of this computational framework for biopharmaceutical discovery against emerging targets. Taken together, these findings support the use of 3D Zernike descriptors for structural clustering of antibodies and epitopes and for guiding antibody repurposing.
Seal, S.; Zalte, A. S.; Araripe, D. A.; Gomes, R. A.; Korani, D.; Shekhar, M.; Siramshetty, V. B.; Patra, A.; Mou, Z.; Yu, X.; Kuhn, D.; Weskamp, N.; Ash, J.; Cheng, A. C.; Fang, C.; Price, D.; Aldeghi, M.; Rodriguez-Perez, R.; Clevert, D.-A.; Engkvist, O.; Deibler, K.; Rouquie, D.; Reutlinger, M.; Richmond, N. J.; Ainsley, J.; Ledeboer, M.; Green, W. H.; Bender, A.; Wognum, C.
Show abstract
Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Sholklapper, T. N.; Li, M.; Srivastava, A.; Wagh, A.; Handorf, E.; Beck, J. R.; Abbosh, P.
Show abstract
Importance There is growing interest to avoid radical cystectomy (RC) in patients with muscle-invasive bladder cancer (MIBC) who receive neoadjuvant chemotherapy and achieve pathological complete response (ypCR). To achieve this goal, molecular biomarkers will likely need to be used to enhance clinical staging given the limitations of evaluation by cystoscopy, cytology, and cross-sectional imaging. There are no studies evaluating whether safe RC avoidance (SaRCA) would be cost effective and what impacts it would have on quality of life (QoL) and survival. Objective This study models the potential economic, QoL, and survival costs/benefits of a ypCR biomarker as it relates to SaRCA using a decision analysis and Markov Model (MM). Methods/Materials A decision tree and MM was created to compare the expected costs of initial treatment, QoL, and survival under one strategy where all patients undergo RC after neoadjuvant treatment versus an alternative strategy where all patients would be subjected to the biomarker test with biomarker-positive patients (those with presumed residual disease) undergoing RC, while biomarker-negative patients (presumed complete responders) would undergo surveillance for up to 20 years. ypCR rates to neoadjuvant therapy, survival with and without RC, quality adjusted life years (QALY), and costs were abstracted from the literature. Test cost, sensitivity, and specificity were also abstracted from the literature for multiple clinical or liquid biopsy approaches. Results Broadly, SaRCA approaches are cost effective with the exception of systematic endoscopic evaluation (SEE). All testing approaches result in higher QALY and overall life expectancy compared to no testing. The cost of the test is offset by decreased usage of RC to realize a cost savings. These domains are further improved when cisplatin-based chemotherapy is replaced with emerging neoadjuvant therapies. Conclusions and relevance Modeling supports the development of accurate biomarker tests which can distinguish residual disease states to enable SaRCA. Such a biomarker could be used to avoid an expensive and risky operation, and unexpectedly would provide a survival benefit by reducing the number of perioperative mortalities in patients achieving ypCR. Development of an accurate biomarker-based test is likely to reduce cost and increase QoL and survival. An accurate biomarker test would have utility for patients, payers, hospitals, and physicians.
Takeuchi, T.; Nomiya, A.
Show abstract
Background: A 2019 report from our institution described a multilayer artificial neural network (ANN) for predicting prostate cancer at biopsy in 334 patients, trained with TensorFlow 1.x and evaluated at three fixed step counts without separating hyperparameter selection from test evaluation. We re-analyzed an expanded cohort from the same institution using contemporary machine-learning practice. Methods: We pooled all available biopsy episodes from the same institutional database (n = 526; 524 after excluding one non-binary outcome code and one record with missing digital rectal examination [DRE] data), retaining the same seven predictors used in the original report (age, prior biopsy history, PSA, prostate volume, DRE, and MRI diffusion-weighted imaging findings in the peripheral and transition zones). Because 27 patients contributed more than one biopsy episode, we used patient-ID-grouped, stratified k-fold cross-validation (StratifiedGroupKFold; scikit-learn 1.8.0) with 3 and 5 folds, repeated over 10 random partitions, to avoid leakage between folds. Four classifiers were compared: L2-regularized logistic regression, gradient boosting, random forest, and a shallow (single hidden layer) multilayer perceptron. Two outcomes were modeled: detection of any prostate cancer, and detection of clinically significant prostate cancer (Gleason score [≥] 7). Results: Any-cancer prevalence was 55.7% (292/524) and Gleason score [≥] 7 prevalence was 39.7% (208/524). With repeated 5-fold cross-validation, gradient boosting gave the highest discrimination for any prostate cancer (mean AUC 0.826, 95% CI 0.823-0.830) and for Gleason score [≥] 7 (mean AUC 0.855, 95% CI 0.852-0.859), closely followed by random forest and logistic regression (AUC 0.81-0.85). The shallow multilayer perceptron performed worse and less consistently than the other three models (any-cancer AUC 0.671; Gleason score [≥] 7 AUC 0.742) and than the deeper five-hidden-layer ANN reported in 2019. Results with 3-fold cross-validation were essentially unchanged. Conclusions: In an expanded cohort, regularized logistic regression, gradient boosting, and random forest all discriminated prostate cancer at biopsy at least as well as the previously reported multilayer ANN, using far simpler models and a methodology that separates hyperparameter tuning from performance estimation. A shallow neural network offered no advantage over these simpler alternatives in this sample size. This is a preprint; the study has not undergone external peer review.
Pusparum, M.; Thas, O.; Ertaylan, G.
Show abstract
Conventional univariate reference intervals (UniRIs) are widely used to identify abnormal biomarker values, but they evaluate each biomarker independently and do not account for coordinated deviations between biomarkers. We developed and evaluated a joint reference region (JRR) framework for plasma proteomics data using the Olink proteomics dataset generated by the UK Biobank Pharma Proteomics Project, covering approximately 3,000 plasma proteins. JRRs were estimated for selected protein pairs in a healthy reference subset, while UniRIs were estimated separately for individual proteins using the nonparametric method. Both approaches were then evaluated in ICD-defined disease subsets. Biomarker discovery revealed sparse and heterogeneous disease--protein associations, with some proteins recurring across multiple phenotypes and others showing more disease-specific patterns. The added value of JRRs varied across diseases and protein pairs. Across evaluated protein pairs, 56.5\% showed higher sensitivity under the JRR framework than the UniRI of the first protein, and 47.3\% showed higher sensitivity than the UniRI of the second protein. At the disease level, the median proportion of protein pairs with improved JRR sensitivity was 0.57. JRRs were most informative when univariate detection was limited but a subset of diseased observations was flagged only by the joint region. These findings suggest that JRRs provide a complementary approach to UniRIs by capturing abnormal joint biomarker configurations in high-dimensional proteomics data.
Polunina, P. V.; Maier, W.; Rubin, A. F.
Show abstract
The evolutionary accessibility of a protein mutation depends on the sequence background in which it arises and its lineage history, yet most protein language models estimate sequence plausibility without explicitly considering the ordered sequence changes through which descendants arise. We developed evoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features. We demonstrated our approach using SARS-CoV-2 spike protein, pairing sequences from early Omicron lineages according to their positions on a mutation-annotated phylogeny, and evaluating model performance on sequence pairs from later lineages. Prompt-masking experiments showed that incorporating phylogenetic context substantially increased reliance on the supplied input information compared with a sequence-only model. Generated descendant sequences accurately reproduced the positional distribution of mutations observed during viral evolution, with strong correlations between predicted and observed mutation-frequency profiles for both the receptor-binding domain (Spearman's {rho} = 0.823) and the full spike protein ({rho} = 0.736). Although prediction accuracy for individual substitutions decreased with increasing evolutionary distance, the model consistently captured aggregate mutational patterns across the spike protein. Model-assigned mutation probabilities were also enriched among substitutions experimentally tolerated in deep mutational scanning assays of Omicron BA.2 receptor-binding domain expression (1.19-fold enrichment) and ACE2 binding (1.04-fold enrichment), despite the model being trained solely on observed ancestor-descendant sequence pairs and associated phylogenetic context features. These results demonstrate that explicitly providing protein language models with phylogenetic context during sequence generation can recover lineage-specific mutational patterns and yields probabilistic predictions consistent with experimentally measured functional constraints. evoPLM-Tree provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.